Papers with measurement theory
Evaluating Readability and Faithfulness of Concept-based Explanations (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for evaluating concepts from different perspectives lack a unified formalization. |
| Approach: | They propose a formal definition of concepts generalizing to diverse concept-based explanations’ settings and apply it to other types of explanations or tasks. |
| Outcome: | Extensive experimental analysis was carried out to determine the evaluation measures for explanation evaluation measures. |
Generative Personality Simulation via Theory-Informed Structured Interview (2026.eacl-long)
Copied to clipboard
Pengda Wang, Huiqi Zou, Han Jiang, Hanjie Chen, Tianjun Sun, Xiaoyuan Yi, Ziang Xiao, Frederick L. Oswald
| Challenge: | Personality structured interviews are often lacking in advancing social science research. |
| Approach: | They propose a method to incorporate psychological insights into LLM simulations . they use a measure theory grounded evaluation procedure to evaluate reliability and validity . |
| Outcome: | The proposed method improves human-like heterogeneity in LLM-simulated personality data and predicts personality-related behavioral outcomes. |
Reliability of Topic Modeling (2025.naacl-long)
Copied to clipboard
| Challenge: | Topic models allow researchers to extract latent factors from text data and use those variables in downstream statistical analyses. |
| Approach: | They propose to use McDonald's as a benchmark to evaluate topic model reliability. |
| Outcome: | The proposed model is based on McDonald's , which provides the best encapsulation of reliability on synthetic and real-world data. |
Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics are conflated and can mislead models, resulting in downstream harms. |
| Approach: | They propose a framework for conceptualizing and evaluating the reliability and validity of evaluation metrics based on empirical data. |
| Outcome: | The proposed framework formalizes the source of measurement error and offers statistical tools for evaluating evaluation metrics based on empirical data. |
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems (2025.findings-acl)
Copied to clipboard
Emma Harvey, Emily Sheng, Su Lin Blodgett, Alexandra Chouldechova, Jean Garcia-Gathright, Alexandra Olteanu, Hanna Wallach
| Challenge: | Existing tools for measuring representational harms caused by large language model systems are not useful for practitioners. |
| Approach: | They examine the extent to which public instruments are used to measure representational harms caused by large language model-based systems. |
| Outcome: | The proposed instruments do not meet the needs of practitioners evaluating large language model-based systems. |